Papers with linguistic processing

8 papers
lingvis.io - A Linguistic Visual Analytics Framework (P19-3)

Copied to clipboard

Challenge: Using a modular framework, linguistic visual analytics applications can be rapidly prototypized using a web-based framework.
Approach: They propose a modular framework for rapid prototyping of linguistic, web-based, visual analytics applications.
Outcome: The proposed framework supports rapid prototyping of linguistic, web-based, visual analytics applications.
Visio-Linguistic Brain Encoding (2022.coling-1)

Copied to clipboard

Challenge: Existing studies have failed to explore co-attentive multi-modal modeling for visual and text reasoning.
Approach: They propose to use image and multi-modal Transformers to reconstruct fMRI brain activity . they use two popular datasets to study visual and text reasoning .
Outcome: The proposed model outperforms existing models on two popular datasets . the results raise the question whether visual processing is affected implicitly by linguistic processing .
BabelDOC: Better Layout-Preserving PDF Translation via Intermediate Representation (2026.acl-demo)

Copied to clipboard

Challenge: Existing document translation pipelines face a tension between linguistic processing and layout preservation.
Approach: They propose a framework for layout-preserving PDF translation that decouples visual layout metadata from semantic content.
Outcome: The proposed framework improves layout fidelity, visual aesthetics, and terminology consistency over representative baselines while maintaining competitive translation precision.
RusConText Benchmark: A Russian Language Evaluation Benchmark for Understanding Context (2025.acl-srw)

Copied to clipboard

Challenge: a new context understanding benchmark is proposed for short-context understanding in Russian . the benchmarks focus on broad reasoning tasks or long-concept comprehension, but are limited in their ability to perceive subtle nuances of context.
Approach: They propose a new benchmark for evaluating short-context understanding in Russian . they propose to use four tasks to assess model performance from a specific perspective .
Outcome: The proposed benchmark is adapted to Russian-language data.
Evaluating Language Tools for Fifteen EU-official Under-resourced Languages (2020.lrec-1)

Copied to clipboard

Challenge: Evaluation of language tools available for 15 EU-official under-resourced languages . evaluation of NERC systems was problematic because of lack of universally or cross-lingually applicable named entities classification scheme.
Approach: They evaluate language tools available for 15 EU-official under-resourced languages . they focus on existing NLP platforms that provide models for under-represented languages - stanton core, nl cube, uDPipe .
Outcome: The evaluation of language tools for 15 under-resourced languages is reproducible . the results are below what was reported in the literature and in some cases even better than the ones reported previously.
Interpretability of Language Models via Task Spaces (2024.acl-long)

Copied to clipboard

Challenge: linguistic interpretability is a method used to assess language models' ability to interpret outputs.
Approach: They propose a method to assess LMs' language conceptualisations by 'similarity probing' and a technique to fine tune them via gradient differentials to disentangle the learning signals of linguistic phenomena.
Outcome: The proposed method generalises larger models to overarching general concepts for linguistic tasks, and the generalisation patterns are stable throughout training and not marked by incisive stages.
FLAG-TRADER: Fusion LLM-Agent with Gradient-based Reinforcement Learning for Financial Trading (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have impressive reasoning capabilities in financial tasks, but struggle with multi-step, goal-oriented scenarios in interactive financial markets.
Approach: They propose a framework that integrates large language models with gradient-driven reinforcement learning (RL) policy optimization.
Outcome: The proposed framework improves performance in trading and other financial domain tasks.
Konidioms Corpus: A Dataset of Idioms in Konkani Language (2024.lrec-main)

Copied to clipboard

Challenge: Konkani is a low-resource language spoken by 2.5 million speakers . idiomatic sense processing is challenging due to the nature of idioms .
Approach: They propose to use crowdsourced idiomatic sentence identification to build a corpus for idioms in the Konkani language.
Outcome: The proposed corpus consists of 6520 sentences written in the Konkani language.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations